Papers with low-latency streaming applications
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper (2024.findings-emnlp)
Copied to clipboard
Iuliia Thorbecke, Juan Pablo Zuluaga Gomez, Esaú Villatoro-tello, Shashi Kumar, Pradeep Rangappa, Sergio Burdisso, Petr Motlicek, Karthik S, Aravind Ganapathiraju
| Challenge: | a recent study shows that training of ASR models with little to no supervised data is challenging. |
| Approach: | They propose a framework to train streaming Transformer-Transducer models with pseudo-labeled (PL) speech from foundational speech models. |
| Outcome: | The proposed framework can be trained from scratch with pseudo-labeled speech from foundational speech models (FSMs) the proposed framework is validated on 6 languages from CommonVoice and proposes multiple heuristics to filter out hallucinated PLs. |